Repository navigation
Conversation
There was a problem hiding this comment.
All reported issues were addressed across 19 files
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
The report and interim-review calls always went to Google, so trying a model on one's own GPU meant editing the URL by hand. They now go to CODETRIAL_GEMINI_REST_BASE when it is set, read from the process environment like INTERVIEW_ROOM_NAME; unset, nothing changes. The live socket is untouched. scripts/gemini-shim.py answers generateContent from llama-server's OpenAI-compatible endpoint. It turns the Gemini schema into the JSON Schema llama.cpp compiles to a grammar, carries function declarations and calls both ways, passes upstream status codes through so the retry rules still see a 503 as a 503, and gives up on llama-server when the caller would. It caps a request that names no output limit, so a model stuck repeating itself ends as MAX_TOKENS instead of out-waiting the shim, and --thinking off turns thinking off for every request, which Gemma 4 needs for a check that names no thinking budget. tests/test_gemini_shim.py covers the mapping and runs in the gate. An ignored unit test makes one real report through the base. Against gemma-4-12b on an RTX 5070 Ti it wrote a valid report in 16 seconds.
A 9B to 14B model on one consumer GPU takes 14 to 32 seconds per report call, against a 20-second attempt limit sized for Gemini, so a local model timed out on most calls. When CODETRIAL_GEMINI_REST_BASE points anywhere but Google, the attempt limit is 45 seconds and the report deadline 250, so five attempts and their doubling backoffs still fit; Gemini keeps 20 and 125. Only the server knows which deadline is in force, so the page's wait before offering to leave and its wait on a regenerated report now come from /runtime-config.js, and the page's own values stay the hosted ones as floors.
A report whose improvement plan called the exercise "a 'Two Sum' style problem" was sent back with only "names the published problem" and the field's path. The model could not tell which words broke the rule, wrote the same sentence twice more, and the report was lost. The repair prompt now names the published title and the scenario's title to use instead. The title stays out of the error itself: that error is the failure note the candidate reads, and the title is what it must not show them. The original prompt already carries the title, so the model learns nothing new from it. Against gemma-4-12b on the Two Sum weak-candidate prompt, 2 of 4 reports passed before this and 4 of 4 after, three of them on the first repair.
When the improvement plan and the feedback disagreed, the repair was told only that they did, and a local 12B model sent back the same plan byte for byte on every repair: told that something in a list of four was wrong, it could not find which. The usual cause is a weakness reworded on its way into the plan, or a repeat standing where an improvement was left out. The repair now names each wrong item by index and each improvement with no item, and when there is one of each, the swap. Improvements are counted once each, the way the validator counts them, and an item without a weakness keeps its index. The note the candidate reads is unchanged. Against gemma-4-12b on the Two Sum report prompt, 0 of 6 reports passed before this and 10 of 10 after; 7 of the 10 broke the plan on their first attempt and were repaired.
The scripted check posted to Google's URL by hand, so it could not follow CODETRIAL_GEMINI_REST_BASE the way the report does. It now builds its URL with gemini_generate_content_url, so pointed at scripts/gemini-shim.py it sends what production sends through the shim the report uses, tools included. The direct route the played candidates added, BEHAVIOR_LOCAL_BASE, stays for talking to an OpenAI-compatible server without the shim. README and docs/development.md describe both. Against gemma-4-12b through the shim with --thinking off, the Two Sum script passed in 9 seconds.
|
I ran one interview by hand with everything local: Setup:
What this PR covers:
Found on the interviewer side, outside this PR:
This is one run on one problem. I will run a few more, including a whiteboard interview, before calling the local report path done. |
The phase judge asks for application/json with no response schema, and the shim passed that on unconstrained. gemma-4-12b then wrapped its answer in a ```json fence, the judge's parser refused it, and in a five-minute interview by hand all 15 judgments were dropped: no REACTO step was recorded from speech, so the last hint stayed withheld after the approach had been stated. Such a request now asks llama-server for a JSON object, which is what Gemini returns for it. The live judge test now follows CODETRIAL_GEMINI_REST_BASE, as the rest of the behaviour check does, and prints each judgment. Against the shim it failed before this change and passes after it.
|
The first run found a bug in this PR, now fixed in 4d63990, and a second run with the same setup confirms the fix. The bug. The phase judge requests Second run, same problem and setup, about seven minutes:
The report this time described the bug correctly (inserting before checking, so 3 matched itself), scored all six REACTO phases, marked the behavioral round as skipped, and did not name the LeetCode title. What the judge still gets wrong with gemma. The format is fixed; some of its judgments are not good:
These are judgment errors by a 12B model, not transport problems, so I have not tried to fix them in the shim. If the check should also hold the quote to the phase it claims, that would belong in |
|
A correction to my first comment, where I wrote "Nothing in the session reached Google". That was not quite true.
Every model call was local in both runs: the report, the phase judge and the pause review went to gemma through the shim, and the live interviewer used local whisper, gemma and Kokoro. Neither CodeTrial's log nor either shim's log shows a request to |
|
I opened #257 to collect results from other models and GPUs. It gives step-by-step instructions and a script that runs the report test three times plus the behaviour check against this branch, along with a results template. My two runs on the 5070 Ti are in it as the reference. The results should show whether the fixed local deadlines (45 s per report call, 12 s per side call) hold on slower hardware, and which models follow the hint and disclosure rules. |
This now carries #99 as well, as asked there. Main had moved a long way since both branches were cut (multi-key reports, doubling backoffs, report recovery, the played-candidate check from #107), so rather than replay 16 commits over it, the work is rebuilt on current
mainas five commits. The old history is still readable on #99.Summary
The report, the quiet-pause reviews and the interviewer behaviour check can now run on a model on the operator's own GPU instead of Gemini. The live interviewer is untouched and still talks to Gemini.
CODETRIAL_GEMINI_REST_BASE, read from the environment likeINTERVIEW_ROOM_NAME, points everygenerateContentcall at another server. Unset, the URL is what it was.scripts/gemini-shim.py(standard library only) answersgenerateContentfrom llama-server's OpenAI-compatible endpoint. It converts the response schema to the JSON Schema llama.cpp compiles to a grammar, carries function declarations and calls both ways, passes upstream status codes through so a 503 still reads as a 503, and gives up on llama-server when the caller would. A request with no output limit gets one, and--thinking offturns thinking off for every request, which Gemma 4 needs for the behaviour check.tests/test_gemini_shim.pyruns in the gate./runtime-config.js, and the page's own values stay the hosted ones as floors.BEHAVIOR_LOCAL_BASEfrom Hold played candidates to the interview rules #107 stays as the direct route, without the shim.Numbers
RTX 5070 Ti, gemma-4-12b Q4_K_M on llama.cpp, shim with
--thinking off:log_hint.One prompt, one machine: read these as "this works end to end", not as a comparison with Gemini.
Test plan
binary_web_gives_a_local_report_base_the_longer_waitstarts the binary with the base set and checks/runtime-config.jssays 260000 and 265; without it, the runtime-config test pins 135000 and 140.report-recovery.test.js: the server can raise the retry wait and cannot lower it.a_local_model_writes_a_reportis an ignored unit test that makes one real report through the base../scripts/test.shpasses locally: 844 JS and 1228 Rust tests, none skipped, formatting clean. Every commit builds clean under clippy.cargo mutants --in-diffon this diff: 42 mutants, 38 caught, 4 unviable, none missed. cargo-audit and shellcheck were not installed here.Not tested: a full interview end to end with the report coming from a local model, and anything against Gemini.